Papers with vision APIs
MM-Reasoner: A Multi-Modal Knowledge-Aware Framework for Knowledge-Based Visual Question Answering (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent knowledge-based visual question answering approaches miss visual information captured by captions and cannot fully utilize the visual information required to answer the question. |
| Approach: | They propose a framework that extracts visual information from an image and prompts an LLM to extract query-specific knowledge from the extracted textual information. |
| Outcome: | Empirical results show that MM-Reasoner achieves state-of-the-art performance on several KVQA datasets. |